Back

Journal of Speech, Language, and Hearing Research

American Speech Language Hearing Association

Preprints posted in the last 30 days, ranked by how well they match Journal of Speech, Language, and Hearing Research's content profile, based on 13 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Evaluating Goodness of Pronunciation and Phonological Posteriors as Objective Markers of Speech Severity in Motor Speech Disorders

Wang, F.; Utianski, R. L.; Duffy, J. R.; Barnard, L. R.; Botha, H.

2026-07-16 neurology 10.64898/2026.07.14.26358076 medRxiv
Top 0.1%
23.3%
Show abstract

This study examined the extent to which goodness of pronunciation (GoP) scores and phonological posterior probabilities capture perceptual ratings of speech severity in individuals with motor speech disorders (MSD). Speech recordings of the word catastrophe were obtained from 489 participants, including 333 neurologically typical controls and 156 individuals with MSD. GoP scores were derived using traditional acoustic features and self-supervised speech representations, including WavLM and XLS-R, across multiple modeling approaches, while phonological posterior probabilities were extracted using Phonet. Model performance was evaluated using Kendall's rank correlations, regression, and receiver operating characteristic analyses against speech-language pathologists' perceptual ratings of sound distortion and intelligibility. Both GoP and phonological posterior probabilities were significantly associated with perceptual ratings. Self-supervised speech representations substantially outperformed traditional acoustic features, with WavLM-based GoP using k-nearest neighbors achieving the strongest performance. Across correlation, regression, and classification analyses, GoP consistently outperformed phonological posterior probabilities for both sound distortion and intelligibility. Age and gender had minimal influence on model-derived measures or their relationships with perceptual ratings. These findings demonstrate the value of self-supervised GoP as an objective measure of speech impairment while highlighting the complementary role of phonological posterior probabilities in characterizing articulatory aspects of motor speech disorders.

2
The Listening Effort Profile of Eye Movements: Easy, Difficult, and Impossible Speech Comprehension

Herrmann, B.; Fink, L. K.; Pandey, P. R.; Johnsrude, I.; Ryan, J. D.

2026-07-03 neuroscience 10.64898/2026.06.30.735702 medRxiv
Top 0.1%
9.5%
Show abstract

Speech comprehension in noisy environments often requires cognitive effort, but listeners may disengage when comprehension becomes impossible. Eye movements have recently emerged as a promising new measure of listening effort, but it remains unclear whether eye movements are sensitive to the full effort profile across easy, difficult, and impossible speech comprehension. Across four experiments, participants listened to sentences at easy, difficult, and impossible levels of multi-talker background babble while pupil size and eye movements were recorded. Pupil size generally followed the expected inverted u-shaped effort profile: low for easy speech, maximal for difficult but still intelligible speech and lower again for impossible speech, although this pattern partly reflected sustained, condition-specific differences and not only sentence-evoked responses. Gaze dispersion - measuring the spread of eye movements - decreased with high temporal selectivity during difficult relative to easy and impossible speech, indicating reduced eye movements during active, effortful listening. However, gaze dispersion was also lower, but less temporally selective, during impossible compared to easy listening, especially in non-baseline-corrected analyses, suggesting that reduced eye movements do not index listening effort uniquely. Instead, eye movements appear to reflect both attentional engagement during difficult listening and disengagement or inward attention when meaningful listening is no longer possible. These findings indicate that pupil size and eye movements provide complementary indices of listening-related cognition, and highlight the integration of listening, cognition, and motor systems.

3
Automated Detection of Motor Speech Disorders and Subtype Classification

Wang, F.; Utianski, R. L.; Barnard, L. R.; Stricker, J. L.; Clark, H. M.; Meade, G. F.; Jones, D. T.; Whitwell, J. L.; Josephs, K. A.; Duffy, J. R.; Botha, H.

2026-07-19 neurology 10.64898/2026.07.16.26358268 medRxiv
Top 0.1%
7.8%
Show abstract

Motor speech disorders (MSDs) are early markers of neurological disease, but expert perceptual analysis is rarely available outside specialized centers. Automated speech analysis offers a scalable alternative, yet prior studies have not systematically compared modeling approaches or assessed clinically relevant metrics in independent datasets. This study compared static acoustic features, articulatory informed Phonet features, and self-supervised pretrained models for binary and multi label MSD classification. We trained and evaluated models on 583 speech samples using speaker level splits. Baseline models included logistic regression and Gated Recurrent Units (GRUs) trained on eGeMAPS and MFCCs. We extracted three types of Phonet derived features and evaluated pretrained HuBERT and SSAST models in frozen, partially fine-tuned, and fully fine-tuned configurations. Binary classification distinguished MSDs from controls, while multi label classification identified six MSD subtypes. Models were assessed using validation AUC, and cut points were tested on two independent datasets. Pretrained and Phonet based models substantially outperformed static acoustic features. In binary classification, HuBERT achieved the highest AUC (0.95), while compact Phonet derived GRUs achieved comparable performance (up to 0.94). These models generalized well to independent datasets, maintaining high sensitivity (0.94) and specificity (0.97). In multi label classification, Phonet models achieved the highest macro average AUC (0.86), but threshold-based subtype performance declined on unseen data. Automated MSD detection is feasible and clinically promising. Binary classification generalized well, whereas multi label classification showed limited threshold stability across datasets.

4
Neurophysiological Evidence for Reduced Use of Prior Sound Patterns to Shape Speech Processing in Autism

Lau, J. C. Y.; McHaney, J. R.; Goldman, L.; Robinshaw, K.; Mou, F.; McFarlane, K.; Chandrasekaran, B.; Losh, M.

2026-07-10 neuroscience 10.64898/2026.07.09.737536 medRxiv
Top 0.1%
3.4%
Show abstract

Reported perceptual differences in autism may arise from reduced use of prior context to shape incoming sensory input. Speech perception provides a critical test of this account because stable perception requires listeners to integrate variable acoustic signals with contextual expectations. This study examined context-dependent modulation of speech encoding in autistic and non-autistic adults using the frequency-following response (FFR), a neurophysiological measure of phase-locked auditory encoding. Participants heard English intonational pitch contours presented in repetitive and variable contexts while EEG was recorded. Principal component analysis of FFR metrics yielded components indexing neural encoding fidelity and timing. Non-autistic participants showed enhanced encoding fidelity in more predictable contexts, whereas autistic participants showed reduced context-dependent modulation. Neural encoding timing also showed divergent context effects across groups, suggesting altered balance between feedback-based predictive mechanisms and locally driven adaptation processes. Within the autistic group, greater context-related modulation of encoding fidelity was associated with lower ADOS-2 Social Affect severity but poorer speech-in-noise perception, suggesting that the functional impact of contextual modulation depends on input reliability and task demands. These findings indicate that context-dependent modulation of speech encoding is altered in autism and may contribute to individual differences in auditory and social-communicative function.

5
Neural Tracking of Speech Envelope as an Index of Spatial Release from Masking

Galeano-Otalvaro, J.-D.; Dieudonne, B.; Francart, T.; Wouters, J.

2026-07-02 neuroscience 10.64898/2026.06.29.734758 medRxiv
Top 0.1%
3.2%
Show abstract

Understanding speech in noisy environments relies strongly on binaural cues such as interaural time differences (ITDs) and interaural level differences (ILDs), which support spatial hearing and the segregation of competing sound sources. When these cues are degraded, listeners experience substantial difficulty in complex acoustic environments. Behavioural measures of binaural benefit, such as binaural masking level differences (BMLDs), binaural intelligibility level differences (BILDs), and spatial release from masking (SRM), are well established in normal-hearing (NH) listeners, but they require an active behavioural response. Neural speech tracking using electroencephalography (EEG) has emerged as a promising approach for quantifying neural processing of continuous speech, yet its sensitivity to spatial hearing cues remains insufficiently characterised. In this study, we investigated the neural correlates of spatial release from masking in NH listeners using EEG-based neural speech tracking. Nineteen participants listened to continuous Dutch speech stories presented with masking noise under two spatial configurations, collocated (S0N0) and spatially separated (S0N90), across multiple signal-to-noise ratios (SNRs). Neural tracking of the speech envelope was quantified using both envelope reconstruction and temporal response function (TRF) analyses. Spatial separation enhanced neural tracking of the target speech envelope, particularly at challenging SNRs where behavioural SRM was also observed. TRF analysis further revealed condition-dependent morphologies, including increased amplitudes and decreased latencies of late cortical components consistent with spatial unmasking effects. These neural differences were most pronounced at low SNRs, where spatial cues provide the greatest perceptual benefit. Together, these findings demonstrate that neural speech tracking captures cortical signatures of spatial unmasking and closely reflects behavioural improvements in speech understanding. Establishing these relationships in NH listeners supports the development of objective neural measures for evaluating binaural benefit in difficult-to-test populations.

6
Neural tracking of stressed syllables in Dutch nursery rhymes relates to vocabulary outcomes in a large, longitudinal sample

Klis, A.;Menn, K.;Cetincelik, M.;Snijders, T.;Junge, C.

2026-06-29 Developmental Biology 10.64898/2026.06.24.734253 medRxiv
Top 0.1%
3.2%
Show abstract

Speech consists of regularities at different timescales. Already during infancy, neural electrophysiological activity aligns to these rhythms. The degree to which infants exhibit neural tracking of speech can be linked to their language development. In this study, we examined how the neural tracking of sung speech develops across age, from infancy to early childhood, and across different frequency bands (i.e., at the stress, syllabic, and phonemic rates), and whether neural tracking at each frequency and age predicts childrens language outcomes. We included 2565 children of the longitudinal YOUth cohort. Children listened to Dutch sung nursery rhymes while EEG was recorded at three measurement waves. After preprocessing the data, we included 955 children at 5 months, 1048 children at 10 months, and 795 children at 2-4 years. The final sample consisted of 750 children who also completed a receptive vocabulary test at 2-4 years. Children from 5 months onwards showed significant neural tracking of stressed syllables, syllables, and phonemes, measured with speech-brain coherence (SBC). Unexpectedly, there were no developmental changes in SBC across different frequency bands from infancy to early childhood. As expected, children with larger receptive vocabularies showed increased SBC in the stressed syllable rate. These findings suggest that stronger tracking of stressed syllables is related to individual differences in language ability.

7
Validity and Reliability of the Novel Indonesian Instrument for Aphasia Diagnosis (IDEA)

Prawiroharjo, P.; Fakhri, A.; Gabrielle, A.; Martalia, V.; Rahmayani, S. A.; Wijaya, V. G.

2026-07-19 neurology 10.64898/2026.07.17.26358303 medRxiv
Top 0.1%
2.9%
Show abstract

Aphasia diagnosis in Indonesia remains challenging due to limited culturally and linguistically appropriate instruments. Widely used tools such as the Boston Diagnostic Aphasia Examination (BDAE) and Western Aphasia Battery (WAB) are not adapted to the Indonesian context, while Tes Afasia untuk Diagnosis, Informasi, dan Rehabilitasi (TADIR) provides screening but lacks diagnostic accuracy. To address this gap, we developed the Instrumen Diagnosis dan Evaluasi Afasia (IDEA) for native Indonesian speakers and evaluated its validity, reliability, and normative cutoff values in cognitively healthy Indonesian adults. Eighty-three cognitively normal adults (screened using MoCA-Ina) with no history of neurological disease were assessed using IDEA, which evaluates six language domains. Items were adapted from existing tools and reviewed by experts. Content validity, internal consistency (Cronbachs alpha), and construct validity (Exploratory Factor Analysis) were analyzed using SPSS v25. A total of 83 participants were included (median age = 55.81 years, 54% secondary education). IDEA demonstrated good feasibility, with an average completion time of 45-60 minutes depending on participant engagement. Content validity was established by unanimous expert consensus. Construct validity showed meritorious sampling adequacy (KMO = .872) and significant sphericity (Bartletts test {chi}^2 (15) = 278.523, p<.001), supporting factor analysis. Internal consistency showed good reliability across six domains (Cronbachs = 0.896). IDEA is a valid and reliable tool for assessing aphasia in Indonesian natives. It is a culturally appropriate assessment tool which offers structured, domain-based evaluation and supports differential diagnosis of both classical and progressive aphasia syndromes. Keywords: Aphasia, Language Assessment, Indonesian, IDEA, Validity

8
Implicit visuomotor adaptation to clamped feedback is reduced in adults who stutter

Liu, J.; Loudermilk, K.; Kim, K. S.

2026-06-29 neuroscience 10.64898/2026.06.24.734039 medRxiv
Top 0.1%
2.4%
Show abstract

It has been demonstrated that people who stutter exhibit atypical motor control not only in speech tasks but also movements in the non-speech effector system, such as finger or arm motion. Notably, studies have reported that people who stutter show limited sensorimotor adaptation (i.e., updating subsequent movements in response to sensory errors) in both speech auditory-motor (i.e., updating speech movements in response to altered auditory feedback) and upper limb visuo-motor (i.e., updating arm movements in response to altered visual feedback) tasks. Given that speech auditory-motor adaptation is mostly if not entirely implicit (i.e., participants are unaware of the learning), it is thought that people who stutter have limited implicit adaptation in the speech effector system. It remains unclear however, whether such limited implicit learning also extends to upper limb visuomotor adaptation. Here, we examined implicit visuomotor learning in adults who stutter through the means of arm reaching adaptation to clamped visual feedback which provides a cursor that is fixed in direction (8{degrees} counterclockwise from targets) regardless of the participants actual hand location. All participants gradually adjusted their reach angle towards the clockwise direction, adapting in response to clamped feedback, but adults who stutter showed less adaptation compared to adults who do not stutter. In addition, computational modeling suggests that this implicit adaptation difficulties in stuttering individuals may reflect reduced error sensitivity. Together, our findings suggest that implicit sensorimotor learning difficulties in adults who stutter may generalize across multiple effector systems, providing important implications for understanding sensorimotor mechanisms underlying stuttering. Significance statementBy employing the clamped visual feedback paradigm during arm reaching movements, we demonstrated that adults who stutter showed less implicit visuomotor adaptation compared to adults who do not stutter. This study provides the first evidence that implicit sensorimotor adaptation limitations in developmental stuttering generalize across multiple effector systems. Our findings not only add to a growing body of evidence that stuttering is associated with domain-general sensorimotor difficulties but also point to specific underlying processes that may lead to stuttering.

9
Older adults show overexaggerated and larger noise-related degradation in their neural tracking of speech

MacLean, J.; Bidelman, G.

2026-07-03 neuroscience 10.64898/2026.07.03.736364 medRxiv
Top 0.1%
2.1%
Show abstract

Background: Speech-in-noise (SIN) perception is a difficult everyday listening task that becomes more difficult with age. Neural tracking of target speech is associated with successful speech perception in clean and noise-degraded listening environments. How aging impacts neural tracking of speech and relates to behavioral decrements in older adults' SIN perception remains unclear. To address these questions, we measured neural speech tracking during a continuous SIN perception task in younger and older adults via multichannel EEG. Method: Participants (n=83) monitored a continuous stream of syllables (~4.5 Hz) presented in quiet and noise conditions during EEG recordings. We assessed neural phase-locking value (PLV) to the acoustic speech envelope to investigate interactions between aging, hearing loss, and stimulus noise on neural synchronization to speech. Results: Compared to younger adults, older adults demonstrated less behavioral sensitivity to noise effects than young adults and had higher overall PLV to target speech. Older adults also showed greater noise-related degradations in neural speech processing relative to younger listeners. Age remained a strong predictor of behavioral responses to speech even after controlling for hearing loss. Covarying for hearing loss removed most age-related effects on neural PLV. Conclusion: Older adults demonstrate overexaggerated neural tracking to ongoing speech presented in quiet and greater noise-related reductions in neurobehavioral speech processing than young adults. Our results support the decline-compensation hypothesis, corroborate unusually large speech envelope encoding in older listeners, and suggest more robust neural synchronization to the speech signal is not always perceptually advantageous.

10
Acoustic and linguistic features of reading reveal early change, progression and function in ataxias

de Belen, R. A. J.; Zheng, Y.; Walsh, M. B.; Hoche, F.; Lin, C.-C.; Stephen, C. D.; Schmahmann, J. D.; White, L.; Belabzioui, H. O.; Kulkarni, D. D.; Patel, S.; Gupta, A. S.

2026-07-14 neurology 10.64898/2026.07.10.26357775 medRxiv
Top 0.1%
1.3%
Show abstract

A major obstacle for clinical trials is the lack of objective, sensitive, and reliable measures that can detect modest changes in disease progression. Here, we determine whether acoustic and linguistic digital speech measures automatically obtained during a functionally relevant passage-reading task capture multiple dimensions of disease in ataxia, including functional communication impairment, subclinical cerebellar dysfunction and disease progression. A total of 157 individuals with ataxia and 84 controls contributed cross-sectional data, and 54 individuals with ataxia and 43 controls contributed longitudinal data within the ongoing Neurobooth natural history study. Participants completed standardized speech recordings, patient-reported outcome measures (PROMs) and neurologist-rated clinical evaluations. A novel speech processing pipeline was developed to automatically transcribe audio recordings, identify word boundaries and extract a predefined set of linguistic and within-word acoustic features. Individuals with ataxia exhibited marked disruption of speech timing, coordination and articulatory control, including slowed speech (d=1.23), prolonged inter-word pauses (d=-0.91), higher/more variable vocal intensity (|d|=0.43-0.51) and altered spectral content (|d|=0.43-0.79) compared to healthy controls. Linguistic features (e.g. speaking rate and within-word pause duration) showed strong associations with clinician-rated severity and PROMs (|r|=0.23-68), indicating alignment with functional communication impairment and patient-perceived disease burden. In contrast, acoustic features derived from cepstral measures captured subtle abnormalities in speech motor control, differentiating not only individuals with ataxia (d=0.65) but also pre-ataxic individuals (d=0.56), and those without clinically evident dysarthria (d=0.45), from controls. These findings indicate that acoustic features reflect subclinical cerebellar motor dysfunction involving impaired temporal coordination and vocal control before overt clinical speech impairment emerges. Longitudinally, several acoustic measures were sensitive to disease progression (MSDR=0.19-0.68), even in cases where clinical scales showed no detectable change. Speech-derived changes correlated with changes in clinical scales and PROMs. Both acoustic and linguistic features exhibited strong intra-session reliability. During passage reading, acoustic and linguistic measures provide complementary but different clinical information in ataxias. Linguistic measures primarily reflect downstream functional consequences of ataxic dysarthria, whereas acoustic measures provide sensitive indicators of subclinical cerebellar motor dysfunction and progression. These findings demonstrate that natural speech analysis can produce digital measures for detecting subclinical disease, quantifying functional impairment, monitoring progression in ataxia, with strong potential for application in clinical trials and remote monitoring.

11
Benchmarking Speech Recognition Models for Medical Consultations in Latin American Spanish: A Comparative Evaluation with Fine-Tuning

Carrillo, R. M.; Carbajal Serrano, A.; Condori Pinedo, P. S.

2026-07-16 public and global health 10.64898/2026.07.14.26358062 medRxiv
Top 0.1%
1.2%
Show abstract

BACKGROUND: Artificial intelligence (AI) medical scribes rely on speech-to-text (STT) models for transcription. Evaluations of STT models in non-English settings remain scarce. We benchmarked ten STT models on medical consultations from Latin American (LatAm) Spanish and assessed whether fine-tuning improves transcription accuracy. METHODS: Ten YouTube videos depicting medical consultations. Human transcriptions were the ground truth. Five open-source models were evaluated: Whisper Large, Whisper Large v3, Whisper Large v3 Turbo, Voxtral Mini 3B, and Canary 1B v2; and so were five close-source models: gpt-4o-transcribe, gpt-4o-mini-transcribe, gemini-2.5-pro, Eleven Labs, and Assembly AI. Whisper Large v3 was fine-tuned. One video was withheld from training. Performance assessed using Word Error Rate (WER), Character Error Rate (CER), BLEU Score, ROUGE-L, BERT Score, and Semantic Similarity on the one withheld video. RESULTS: None of the fine-tuning iterations outperformed the vanilla Whisper Large v3. With the withheld video, Gemini-2.5-pro was the close-source model with the best performance in four of six metrics. In comparison to the close-source models, the fine-tuned model never outperformed the other models (withheld video); conversely, in comparison to the close-source models, the fine-tuned model showed better performance across metrics, for instance: BLEU score (63% vs to 58% for the second-ranking model), BERT (89% vs to 86%), and semantic similarity (89% vs to 83%), CER (19% vs 20%). CONCLUSIONS: Whisper Large v3 and its fine-tuned variant are the best open-source STT models for transcribing medical conversations in LatAm Spanish. These findings provide an evidence base for developing AI medical scribes tailored to Spanish-speaking LatAm.

12
Preserved Spontaneous Interpersonal Entrainment during Rhythmic Synchronization in Autism Spectrum Disorder

Denis, M.; Rosso, M.; Da Fonseca, D.; Schön, D.

2026-07-14 neuroscience 10.64898/2026.07.09.737537 medRxiv
Top 0.2%
1.1%
Show abstract

PurposeInterpersonal entrainment, defined as the tendency of interacting individuals to temporally align their behaviours, is considered a key mechanism supporting social interactions through predictive and multisensory processes. Autism Spectrum Disorder (ASD), characterized by social and communication difficulties, has been associated with atypical predictive and multisensory integration processes, which may affect spontaneous entrainment to others actions during rhythmic interactions. MethodsThe present study investigated spontaneous interpersonal entrainment in 24 autistic and 22 neurotypical young adults using a unidirectional adaptation of the drifting metronomes paradigm developed by Rosso et al. (2021). Participants synchronized their finger tapping with an auditory metronome while either seeing or not seeing a partners hand movements performing the same task at a slightly different tempo. Individual synchronization performance was assessed using asynchrony measures and computational modelling of sensorimotor synchronization, while interpersonal coordination dynamics were quantified using joint recurrence analysis. ResultsVisual exposure to the partners hand movements significantly increased tapping variability while simultaneously enhancing spontaneous interpersonal entrainment. Contrary to previous findings, these effects were comparable across groups, suggesting similar sensitivity to spontaneous low-level interpersonal coupling under stable and predictable conditions. ConclusionOverall, the findings support accounts proposing selective rather than generalized atypicalities in predictive processing and interpersonal entrainment in ASD.

13
Speech clarity shapes auditory attention and visual-signal coupling during multimodal sentence comprehension

Husta, C.; Seijdel, N.; Drijvers, L.

2026-07-14 neuroscience 10.64898/2026.07.13.738151 medRxiv
Top 0.2%
1.1%
Show abstract

Face-to-face communication requires listeners to attend, integrate, and weigh multiple communicative signals, including auditory speech, mouth movements, and co-speech gestures. The contribution of these signals may depend on the reliability of auditory input and the informativeness of the available signals. We utilized rapid invisible frequency tagging (RIFT) with EEG to examine how participants attend to and integrate these different signals in clear and adverse listening conditions. Participants watched videos of an actress producing clear or noise-vocoded sentences. Auditory speech was amplitude-modulated at 58Hz, while the luminance of the gesture and mouth regions was frequency-tagged at 63Hz and 65Hz. Degraded speech elicited stronger responses at the auditory tagged frequency, suggesting increased attentional gain to the auditory signal when listening was challenging. In contrast, clear speech elicited stronger responses at the gesture tagged frequency and a stronger 2Hz intermodulation response (65-63Hz), reflecting enhanced nonlinear coupling between mouth movements and gestures. Finally, in degraded speech, the informativeness of mouth movement, but not gesture, was associated with intermodulation strength, suggesting that the informativeness of mouth movements plays a greater role in multisensory interaction when listening is challenging. Our findings demonstrate that both signal reliability and informativeness shape multisensory integration during spoken language comprehension.

14
Transcranial Photobiomodulation on Language and Cognitive Performance in Down Syndrome: A Pilot Randomized Sham-Controlled Trial

Luchese, F.; Velidi, P. S.; Jaoude, L. B.; Sidelinger, L.; Gersten, M.; Ferreras, B.; Lohmann, C.; Puerto, A.; Clancy, J. A.; McEachern, K.; Sylvester, K.; Chan, S.-t.; Pulsifer, M.; Naeser, M.; Saltmarche, A.; Corcoran, E.; Thurman, A. J.; Abbeduto, L.; Skotko, B.; Cassano, P.

2026-07-06 psychiatry and clinical psychology 10.64898/2026.07.03.26357051 medRxiv
Top 0.2%
1.0%
Show abstract

Background: Down syndrome (DS) is associated with persistent language and cognitive impairments and with abnormalities in cortical oscillatory activity, including in the gamma range. Transcranial photobiomodulation (tPBM) is a noninvasive neuromodulatory intervention with potential benefits for cortical physiology, language, and cognition. Methods: We conducted a pilot randomized, double-blind, sham-controlled trial of repeated 40-Hz near-infrared tPBM in adolescents and young adults with DS. Fourteen participants were randomized 1:1 to active tPBM or sham and received 18 sessions over 6 weeks, followed by short-term and long-term follow-up. Outcomes included resting-state EEG gamma power, connected-speech measures, language and cognitive indices, and selected computerized tasks. Results: Active tPBM did not significantly increase global resting-state EEG gamma power relative to sham at either follow-up. Pre-registered analyses did not show broad treatment benefit across outcomes, although they did identify a significant short-term advantage for active tPBM on grammatical morpheme accuracy in connected speech; in contrast, picture naming favored sham at long-term follow-up. Exploratory mechanistic analyses did not show a robust biological treatment signal. Both active and sham procedures were well tolerated, with no serious adverse events. Conclusions: In this underpowered pilot sample, 6 weeks of 40-Hz near-infrared tPBM--delivered unilaterally, at low power, on target areas--did not demonstrate a meaningful effect size in DS. A dose-finding study for tPBM in DS, also accounting for age of participants, is recommended.

15
Towards a Framework for Case Identification in Pharmacovigilance: Not All Reports are Created Equal.

Fusaroli, M.; Felix China, J.; Sartori, D.; Giunchi, V.; Harmark, L.; Scholl, J.; van Hunsel, F.; Noren, G. N.; Ellenius, J.

2026-07-01 pharmacology and therapeutics 10.64898/2026.06.23.26354546 medRxiv
Top 0.2%
0.9%
Show abstract

Background: Retrieval of adverse event reports based on coded drug-event co-occurrence enables large-scale pharmacovigilance analyses, but yields candidate reports rather than validated cases, risking misinterpretation if used alone. Aim: To develop and apply a framework for identification and characterization of clinically meaningful case series in pharmacovigilance. Methods: We conducted two case studies. The first developed and refined the framework in an information-rich setting, focusing on drug-induced impulsivity across selected drugs; the second tested its applicability in a more routine, information-poor setting, focusing on drug-induced suicidality. Results: In Case 1, non-relevant reports were frequent for drugs with uncertain evidence and negative controls ({approx}20-40%) compared to drugs with established causal roles (4%). The emerging framework assessed relevance based on exposure, event, drug-event relationship, and population. For suspected adverse drug reactions, relevant reports were further characterized by reporter suspicion and evidentiary qualifiers supporting or refuting causality; higher suspicion was associated with more supportive qualifiers. Applied to Case 2, the framework ruled out 69% of reports as non-relevant but highlighted substantial non-assessability (17%). Conclusions: In pharmacovigilance, retrieval is not equivalent to case identification. Relevance is question-specific and shaped by how reports are captured, processed, and retrieved. This can be especially critical for emerging or bias-prone safety questions. Transparent and reproducible case definition and adjudication are essential for interpretable analyses.

16
Sequential Word Properties in Verbal Fluency: Detecting High-Proficiency Cognitive Impairment

Chang, Y.-N.; Wang, Y.-H.; Chou, C.-J.; Liu, Y.-C.; Lambon Ralph, M. A.

2026-07-09 neurology 10.64898/2026.07.06.26357360 medRxiv
Top 0.2%
0.6%
Show abstract

Verbal fluency (VF) tasks are widely used to differentiate patients with cognitive impairment from healthy controls, but total word count produced during these tasks becomes unreliable when patients and controls exhibit comparable proficiency. This study examined, in detail, whether item-level and sequential properties of words produced during a VF task could reliably differentiate high-proficiency patients indistinguishable from controls by word count alone. Seventy-seven native Mandarin Chinese speakers (38 controls and 39 patients with mild cognitive impairment or mild dementia) completed a semantic VF task. Participants were subdivided by proficiency into four groups: high-proficiency controls (HC), low-proficiency controls (LC), high-proficiency patients (HP), and low-proficiency patients (LP). The LC and HP subgroups were matched on semantic fluency scores and thus provided a key focus for the investigation. We examined item-level properties (word frequency, contextual diversity, semantic diversity, surprisal) and sequential properties (positional frequency variation) of the words produced. Significant group differences emerged across item-level psycholinguistic properties, though these were primarily driven by the LP group, with no reliable differentiation between LC and HP. Crucially, positional frequency variation distinguished LC from HP. LC participants began their lists with high-frequency words followed by a systematic decline, whereas HP patients produced words within a consistently narrow frequency band throughout. These findings indicate that item-level psycholinguistic properties alone are insufficient to differentiate HP from LC, whereas sequential word frequency variation provides a potential index of cognitive impairment, reflecting underlying differences in semantic retrieval and memory organisation. Future work with larger samples is needed to validate generalisability.

17
Perspectives in conducting task-based research in pediatric surgical epilepsy patients

Leisawitz, J. P.; Georges, S. F.; Field, A. M.; Asghar, S.; Foox, G.; Watrous, A. J.; Weiner, H. L.; Anderson, A. E.; Hamilton, L. S.

2026-07-08 neuroscience 10.64898/2026.07.02.734030 medRxiv
Top 0.2%
0.6%
Show abstract

Objective: Pediatric epilepsy patients undergoing stereo-electroencephalography (sEEG) for ictal onset evaluation provide a rare window to study the developing brain. While methodological frameworks for task-based sEEG research are well-established in adults, pediatric-specific guidance remains underdeveloped. Furthermore, many pediatric epilepsy patients have comorbidities that might typically exclude them from participating in research. We examine factors that influence research participation and discuss considerations for conducting sEEG research in children. Methods: Here, we present a retrospective analysis of task-based research participation patterns from an NIH-funded study of speech and language representations (1R01DC018579) in 66 patients (ages 4-24) undergoing sEEG monitoring at Texas Children's Hospital to determine whether specific comorbidities influenced research participation. Results: Eighty-nine percent (n=66) of patients approached for consent agreed to participate in the study. Despite high rates of comorbidities including neurocognitive disorder (66.67%), language delay (31.75%), global developmental delay (23.81%), mood disorders (33.33%), ADHD (46.03%), autism spectrum disorder (14.29%) or other cognitive/intellectual disabilities (36.51%), all participants engaged in at least one task. While the majority of these diagnoses did not appear to influence subject participation, global developmental delay was associated with a significant reduction in time spent on active tasks. Discussion: Despite high prevalence of neuropsychological comorbidities among participants, our evidence suggests that these participants contribute meaningfully to studies investigating important developmental questions. We suggest strategies for tailoring task-based research to accommodate the unique needs of individuals in this population. Such practices are important for ensuring that research studies reflect the true diversity of the population.

18
Perceptual consistency in phoneme categorization is driven by neural consistency and predicts improved speech-in-noise performance

Rizzi, R.; Stirn, J. R.; Eisenhut, Z.; Bidelman, G. M.

2026-07-03 neuroscience 10.64898/2026.07.02.736174 medRxiv
Top 0.2%
0.6%
Show abstract

Listeners discretize the speech signal by assigning sounds to phonetic categories, though there is variability in how individuals accomplish categorization. Having more consistent categorization of sounds may be advantageous for understanding speech-in-noise (SIN). Though, it is unclear how different levels of neural processing in the auditory system reflect these perceptual differences. We recorded brainstem frequency-following responses (FFRs) and cortical event-related potentials (ERPs) while listeners actively labeled vowels along an acoustic-phonetic continuum using a visual analog scale. We computed intertrial consistency of neural responses to index the stability of listeners' neural speech representations across stimulus presentations. We also assessed how faithfully midbrain and cortical responses represented stimulus acoustics using representational dissimilarity matrices (RDMs) computed across all token pairs. Neural RDMs were then compared with acoustic and phonetic category RDMs to assess whether FFRs and ERPs carried gradient vs. categorical information of the speech signal. We found greater behavioral consistency during phoneme labeling was correlated with improved SIN scores. Neurally, we found greater cortical or subcortical consistency predicted greater behavioral consistency. RDMs revealed subcortical responses retained more acoustic details, while cortical responses more closely reflected abstract phoneme categories. Our findings reveal important benefits of perceptual consistency to other domains of speech perception. We find perceptual consistency is driven by more consistent encoding of speech at either a cortical or subcortical level. More consistent sensory processing could provide a more stable readout of the speech signal to higher cortical brain areas which could confer advantages to later perceptual processes downstream.

19
Articulatory timing and form support distinct neural benefits during audiovisual speech

Nidiffer, A.; O'Sullivan, A.; Lalor, E. C.

2026-07-15 neuroscience 10.64898/2026.07.14.738583 medRxiv
Top 0.3%
0.6%
Show abstract

In noisy environments, visible speech articulations improve listening comprehension. The benefit derives from several sources, including articulatory timing and shape. Recent research has shown that visual cortex encodes a categorical representation of articulatory features and that visual speech can benefit both acoustic and phonetic feature processing separately. The present study advances the hypothesis that the shape of the articulators specifically influences the categorization of auditory speech in terms of its phonetic features. We tested this by linearly modeling electroencephalographic responses to natural, continuous speech (in noise) in terms of the acoustic and articulatory features of the speech. We compared the performance of these models in conditions where the speech was accompanied by a natural video of the speaker with their mouth visible, and a video where their mouth was covered by a dynamic ellipse obscuring articulatory shape but preserving dynamics. The dynamic mask reduced comprehension, neural processing of phonetic features, the associated multisensory benefits, and indices of visual-only linguistic processing over occipital scalp. Our findings support substantial visual involvement in speech comprehension, derived largely from the shape of the articulators. They also corroborate several proposals involving audiovisual speech processing hierarchy and the nature of the information contained in visible speech. HighlightsO_LIVisual speech provides at least two forms of information to enhance acoustic speech processing: redundant temporal dynamics and complementary articulatory information C_LIO_LICovering the mouth with a dynamic mask preserves horizontal and vertical lip movement information, but largely removes articulatory detail C_LIO_LIVisual speech with a mask preserves some general multisensory benefits but removes visual linguistic information and its ability to enhance auditory processing at the level of phonetic features. C_LI

20
Pitch motor areas contribute to the perception of prosodic categories in speech

BAEK, S.-C.; Kim, S.-G.; Maess, B.; Grigutsch, M.; Sammler, D.

2026-06-26 neuroscience 10.64898/2026.06.22.733802 medRxiv
Top 0.3%
0.6%
Show abstract

Prosody is a fundamental aspect of speech characterized by suprasegmental features such as pitch. Prosodic pitch contours are used to convey speakers intentions, for example, to make a statement or ask a question. Understanding these intentions requires abstracting continuous, variable pitch information into discrete categories. Category perception has been proposed to recruit the motor system in an effector-specific manner, whereby cortical areas controlling motor effectors support speech sound recognition by identifying articulatory gestures. However, it remains unclear whether effectors involved in pitch production similarly contribute to prosodic category perception. To address this question, we collected magnetoencephalography data from 29 participants (15 females) while they first sang pitches arranged in five-tone melodies and then identified the prosody (Statement vs. Question) of single words varying in pitch contour along a five-level continuum. Using a region-restricted searchlight approach to decode singing from rest, we localized two premotor regions for pitch production, corresponding to the ventral and dorsal laryngeal motor cortex (LMC). A separate neural decoding analysis revealed that perceived prosodic categories were decodable in these regions, especially from the dorsal LMC that is more closely associated with pitch regulation. Importantly, decoding performance mirrored behavioral discriminability of prosodic categories across the continuum, suggesting that these regions are involved in perceptual decision-making. Finally, pitch motor areas exchanged category-related information with auditory regions, indicating these areas do not merely echo the processing in auditory regions. Together, these findings highlight effector-specific motor support for prosodic category perception, thereby broadening our understanding of motor involvement in speech perception. Significance StatementSpeech perception has been proposed to recruit the premotor cortex, with different subregions linking speech sounds to the articulatory gestures used to produce them. We investigated this idea through prosody--pitch changes in speech conveying meanings such as statements and questions. Using magnetoencephalography, we identified pitch motor areas during a singing task and tested whether they represent perceived prosodic categories. We found that prosodic categories were distinguishable in these regions and that this neural discriminability mirrored behavioral discriminability across clear and ambiguous prosody, suggesting involvement in perceptual decision-making. These findings are unlikely to reflect passive echoes from auditory regions, as pitch motor areas actively influenced them during categorical processing. Our results highlight effector-specific motor support for forming abstract prosodic representations.